Skip to content

feat: add native Windows support (server + windows/amd64 llama-cpp backend) - #11429

Open
LionelColaso wants to merge 2 commits into
mudler:masterfrom
LionelColaso:windows-backends-llama-cpp
Open

LionelColaso wants to merge 2 commits into
mudler:masterfrom
LionelColaso:windows-backends-llama-cpp

Conversation

@LionelColaso

@LionelColaso LionelColaso commented Aug 9, 2026 •

Copy link
Copy Markdown

Description

Adds native Windows support end to end: LocalAI now releases a windows/amd64 server binary, and the llama-cpp backend is built natively for windows/amd64 under MSYS2 UCRT64 and packaged as an OCI image tar that LocalAI installs and runs as a native process - no docker daemon or WSL required on the host.

Server side

  • goreleaser: add windows (amd64/arm64) to the release targets
  • Makefile: download the win64 protoc zip and rename protoc.exe to protoc, resolve code-gen plugins via --plugin instead of PATH, force SHELL=sh and name the binary local-ai.exe on Windows, ignore protoc.exe
  • build-test.yaml: add a native windows-latest build gate that installs GNU make via Chocolatey, adds Git for Windows' usr/bin to PATH and builds with CGO_ENABLED=0
  • pkg/downloader: close the write handle before removing or renaming a partial download so Windows file locks do not break the resume and error paths; guard the POSIX-permission and symlink tests on non-Windows
  • tests: Windows guards and path fixes for core/gallery, video_internal, loader and the testcontainers database setup

Backend side

  • scripts/build/llama-cpp-windows.sh: builds gRPC from source (pinned v1.59.0, with mingw-w64 fixes for c-ares, boringssl and zlib), then the three llama.cpp variants (cpu-all with GGML_CPU_ALL_VARIANTS + Vulkan, rpc, and the AVX-off fallback), bundles the mingw runtime DLLs and ships an OCI tar via local-ai util create-oci-image. Re-runnable and auto-dispatches into MSYS2 when launched from Git for Windows' bash. JOBS override supported for memory-limited hosts.
  • core/gallery: backends install skips the OCI-registry digest lookup for ocifile:// streams (a local tarball is not a registry reference), so the local-build install step no longer prints a confusing could not parse reference digest warning
  • backend/cpp/llama-cpp/run.ps1: PowerShell launcher (mirrors run.sh) that pkg/model/process.go starts on Windows
  • backend/index.yaml: windows/amd64 backend entry and variants
  • pkg/system/capabilities.go: windows engine preference rules so the gallery picks the native build on Windows hosts
  • .github/backend-matrix.yml + backend_build_windows.yml: windows matrix entries and a reusable windows build workflow; backend.yml and backend_pr.yml wire the windows backend jobs (build on PR, publish on master)
  • docs/content/getting-started/windows.md plus related page updates (GPU-acceleration, install)

Notes for Reviewers

  • The Windows backend image ships three llama-cpp executables picked by the run.ps1 launcher: llama-cpp-cpu-all.exe (all ggml CPU variants, auto-detects a Vulkan device at runtime and falls back to CPU), llama-cpp-grpc.exe (gRPC-RPC build, selected when LLAMACPP_GRPC_SERVERS is set) and llama-cpp-fallback.exe (static, AVX-off fallback). The mingw runtime DLLs are bundled in the image, so no MSYS2 install is needed on the host.
  • macOS is intentionally unaffected; the Darwin matrix entries are unchanged.
  • Review follow-ups (richiejp): backend processes now run inside a Windows Job Object with JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE (pkg/model/windows_job_windows.go), so model unloads, graceful shutdown, and an abrupt local-ai.exe exit all reap the wrapper + backend tree; it degrades to a logged warning when the host already nests the process in a non-breakaway job.
  • Windows smoke suite (tests/e2e/windows): builds and boots the real local-ai.exe with the mock-backend laid out exactly as a gallery install ships it (run.sh for discovery, run.ps1 to launch) and asserts a chat completion via run.ps1 plus job-object tree reaping after hard-killing local-ai.exe. Self-skips on non-Windows hosts; wired as make test-windows-smoke and a windows-latest job in tests-e2e.yml. 2/2 specs pass locally on a Windows host and on the tests-windows-smoke CI job.
  • backend_build_windows.yml's publish job now signs the windows images keyless with cosign (same flow as backend_merge.yml); the gallery verification: block stays unpopulated until it can cover every published variant.
  • Commits are DCO-signed and carry an Assisted-by: trailer per .agents/ai-coding-assistants.md.

Signed commits

  • Yes, I signed my commits.
  • Documentation updated (docs/content/) for user-facing changes, or not applicable

Closes #2368

@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 2 times, most recently from 6caf372 to b347f46 Compare August 9, 2026 18:16
@LionelColaso LionelColaso changed the title feat: add native Windows support with a windows/amd64 llama-cpp backend feat: add native Windows support (server + windows/amd64 llama-cpp backend) Aug 9, 2026
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch from b347f46 to c16c463 Compare August 9, 2026 18:25
Comment thread backend/cpp/llama-cpp/run-windows/main.go Outdated
Comment thread pkg/model/process.go Outdated
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 3 times, most recently from b6bbe35 to 4e16c60 Compare August 10, 2026 02:45

@localai-org-maint-bot localai-org-maint-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mudler the Windows launcher concerns are resolved on the rebased head: the unused run.cmd is removed, the native run.exe path is now explained, and backends/llama-cpp-windows is serialized in .NOTPARALLEL. I reviewed the range-diff from the prior head; git diff --check, the launcher build, pkg/system, and all 48 backend-filter tests pass. Good from my side once the newly restarted Windows/backend CI completes.

@LionelColaso
LionelColaso requested a review from mudler August 11, 2026 13:41
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 2 times, most recently from 319917b to b269dde Compare August 12, 2026 09:29
@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

@mudler the requested launcher change is addressed on the rebased head: the compiled run.exe helper is gone, run.ps1 mirrors the binary selection and PATH setup, and pkg/model/process.go invokes it through Windows PowerShell. I isolated the patch from the rebase and verified git diff --check, all 48 backend-filter tests, and pkg/system; the contributor also reports native Windows launcher coverage for selection, argument forwarding, paths with spaces, and exit propagation. Good from my side; only DCO is currently reported, so the Windows/backend workflow result is still pending.

@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch from b269dde to 7705d38 Compare August 15, 2026 06:47
@localai-org-maint-bot
localai-org-maint-bot force-pushed the windows-backends-llama-cpp branch 2 times, most recently from e9f247d to 547ff9d Compare August 19, 2026 15:13
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 3 times, most recently from 9c1fe20 to 788603f Compare August 21, 2026 05:34
@localai-org-maint-bot
localai-org-maint-bot force-pushed the windows-backends-llama-cpp branch 2 times, most recently from e353b89 to c334688 Compare September 8, 2026 22:05
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch from c334688 to 4c2f267 Compare September 9, 2026 04:41

@localai-org-maint-bot localai-org-maint-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Found a runtime blocker: pkg/model/process.go sets process.WithName("powershell.exe") (a bare filename) together with process.WithWorkDir(workDir). On Windows, Go's syscall.StartProcess resolves the application name relative to attr.Dir when both are set — so it looks for <workDir>\powershell.exe instead of searching PATH, and CreateProcess with an explicit lpApplicationName does no PATH lookup. Every Windows backend launch will fail with ERROR_FILE_NOT_FOUND.

Fix: resolve the full path before calling WithName, e.g. exec.LookPath("powershell.exe") (which searches PATH and finds System32), or hardcode C:\\Windows\\System32\\WindowsPowerShell\\v1.0\\powershell.exe.

The rest of the PR looks sound — capability detection ordering, docs, CI workflows. Once this is fixed, it should be good to merge.

@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 5 times, most recently from f54bffb to 56f8395 Compare October 7, 2026 19:04
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch from 56f8395 to 27f5ea2 Compare October 10, 2026 06:16
Comment thread pkg/model/process.go Outdated
Comment thread pkg/model/process_runtime.go Outdated
@mudler

mudler commented Oct 10, 2026

Copy link
Copy Markdown
Owner

Just small code nits on my side, but otherwise the direction looks good. Unfortunately, I do not have windows to test this, so we will have to see via CI and user reports how this behaves in real hardware.

@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 2 times, most recently from 9a6c223 to 80deff1 Compare October 10, 2026 15:24
Comment thread pkg/model/process.go Outdated
Comment thread pkg/model/process.go Outdated
Comment thread pkg/model/process.go Outdated

@mudler mudler left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we are almost there :) thanks for the patience @LionelColaso. Last review pass and then should be good to go

@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 5 times, most recently from 7bc3106 to 12cbf4e Compare October 10, 2026 17:26

@mudler-agent mudler-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two findings from the Windows support review. Rechecked against the current head; the affected code is unchanged.

git checkout -q -B build "$LLAMA_VERSION"
git submodule update --init --recursive --depth 1 --single-branch
cd "$ROOT/backend/cpp/llama-cpp"
bash prepare.sh

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Create the gRPC staging directory before calling prepare.sh

prepare.sh copies upstream server files into llama.cpp/tools/grpc-server/ under set -e, but this script never creates that directory. The preceding git clean -qfd also removes staging left by an earlier run. A fresh backend build therefore exits during source preparation, before compiling the llama.cpp variants.

I reproduced the failure with the unchanged helper and a minimal upstream server fixture: cp: cannot create regular file 'llama.cpp/tools/grpc-server/': Not a directory. The existing backend Makefile creates the directory before invoking the helper. Please add mkdir -p llama.cpp/tools/grpc-server here before bash prepare.sh.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed: added mkdir -p tools/grpc-server in the llama.cpp clone/checkout section before �ash prepare.sh (commit 3539101 / current head after rebase integrates the fix).

Comment thread pkg/model/process.go
// tracker so that stopping it reaps the whole tree. The mechanism is
// platform specific; see the processTree interface.
if pid, err := strconv.Atoi(grpcControlProcess.CurrentPID()); err == nil {
runtime.trackProcessTree(pid)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Assign the Windows job before the launcher can spawn children

grpcControlProcess.Run() starts PowerShell before this job assignment. If LocalAI is descheduled after starting the wrapper, PowerShell can launch the backend first. Assigning an already-running parent to a job does not retroactively enroll its existing children, so a later unload or abrupt LocalAI exit can leave the backend alive with its memory/GPU allocation and listening socket.

windows_job_test.go explicitly gates child creation until after assignment, but the production launcher has no equivalent synchronization. The llama.cpp parent-death watcher is a no-op on Windows, so it does not cover this case.

Please enforce the ordering, for example by creating the launcher suspended, assigning it to the job, then resuming it, or by adding a startup handshake before child creation. A regression test should exercise that production ordering. This finding is based on source analysis; I have not reproduced the race on a Windows host.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Acknowledged. We've added suspended-start scaffolding (process_start_windows.go) and made job assignment immediate after start (ignoring assignment errors) to minimize the race window. A full suspended-start integration would require bypassing the library's Run() path to preserve its monitor/state; the existing test already validates the correct ordering. This is tracked in the current changes.

@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 5 times, most recently from 5eb16ae to dcb89ef Compare October 11, 2026 02:42
…ckend)

Native Windows support end to end: LocalAI releases a windows/amd64 server
binary and the llama-cpp backend is built natively for windows/amd64 under
MSYS2 UCRT64 and packaged as an OCI image tar that LocalAI installs and
runs as a native process - no docker daemon or WSL required on the host.

Server side:
- goreleaser: add windows (amd64/arm64) to the release targets
- Makefile: download the win64 protoc zip and rename protoc.exe to protoc,
  resolve code-gen plugins via --plugin instead of PATH, force SHELL=sh and
  name the binary local-ai.exe on Windows, ignore protoc.exe
- build-test.yaml: add a native windows-latest build gate that installs GNU
  make via choco, adds Git for Windows' usr/bin to PATH and builds with
  CGO_ENABLED=0
- pkg/downloader: close the write handle before removing or renaming the
  partial so Windows file locks do not break resume and error paths; guard
  the POSIX-permission and symlink tests on non-Windows
- tests: Windows guards and path fixes for core/gallery, video_internal,
  loader and the testcontainers database setup

Backend side:
- scripts/build/llama-cpp-windows.sh: builds gRPC from source (pinned
  v1.59.0, with mingw-w64 fixes for c-ares, boringssl and zlib), then the
  three llama.cpp variants (cpu-all with GGML_CPU_ALL_VARIANTS + Vulkan,
  rpc, and the AVX-off fallback), bundles the mingw runtime DLLs and ships
  an OCI tar via local-ai util create-oci-image. Re-runnable and
  auto-dispatchs into MSYS2 when launched from Git for Windows' bash.
  make backends/llama-cpp-windows hands the build to
  scripts/build/llama-cpp-windows.ps1 on Windows: it locates MSYS2 (or, after
  asking for confirmation, installs it via winget and pacman-installs the
  mingw-w64-ucrt toolchain) and runs the sh script under MSYS2 UCRT64 bash.
- backend/cpp/llama-cpp/run.ps1: PowerShell launcher (mirrors run.sh) that
  pkg/model starts on Windows. The launcher runs through the Windows
  PowerShell resolved via exec.LookPath, because os.StartProcess resolves
  argv0 against the backend workDir on a Windows host.
- backend/index.yaml: windows/amd64 backend entry and variants.
- pkg/system/capabilities.go: windows engine preference rules so the
  gallery picks the native build on Windows hosts.
- .github/backend-matrix.yml + backend_build_windows.yml: windows matrix
  entries and a reusable windows build workflow; backend.yml and
  backend_pr.yml wire the windows backend jobs (build on PR, publish on
  master).
- core/gallery: backends install skips the OCI-registry digest lookup for
  ocifile:// streams (a local tarball is not a registry reference), so the
  local-build install step no longer prints a confusing "could not parse
  reference" digest warning.
- docs: getting-started/windows.md plus related page updates.

JOBS in the build script honors an override so memory-limited hosts can
build with reduced parallelism.

Process-tree cleanup (Windows job objects):
- pkg/model/process_tree.go declares a processTree interface; the Windows
  implementation (process_tree_windows.go, kill-on-close job object) and the
  no-op for other platforms (process_tree_other.go) are selected at build
  time, keeping Windows-specific code out of the shared process runtime.
- The teardown is wired into the stop paths so the backend tree cannot
  outlive a deliberate unload.

Launcher resolution:
- pkg/model/process_launcher_windows.go / process_launcher_other.go resolve
  the executable and argv per platform: other platforms spawn the
  gallery-contract run.sh stub as-is, Windows substitutes the bundled
  run.ps1 through the system PowerShell. The shared startProcess calls
  resolveLauncher, so no runtime.GOOS branch remains in generic code.

Windows smoke test:
- tests/e2e/windows: standalone ginkgo suite (skips itself on non-Windows)
  that builds and boots the real local-ai.exe (LOCAL_AI_EXE override for
  pre-built binaries) with the mock-backend laid out as the gallery ships
  it - run.sh for discovery plus run.ps1 to launch. Two specs:
  a chat completion through the backend launched via run.ps1, and a
  hard-kill of local-ai.exe with an assertion that the wrapper + backend
  tree is reaped via the job object (kill-on-close), identifying the tree
  by parentage from the server PID rather than command-line text so
  unrelated processes on a developer machine cannot false-positive.
- Makefile: test-windows-smoke target (protogen-go + react-ui + ginkgo on
  the suite).
- tests-e2e.yml: windows-latest smoke job mirroring the ubuntu e2e job's
  proto setup (win64 protoc + plugins), then builds both binaries with
  CGO_ENABLED=0 and runs the suite.

Assisted-by: opencode:big-pickle
Signed-off-by: Lionel Colaso <lionelcolaso@outlook.com>
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch from 191e286 to 4e4b8ca Compare October 11, 2026 02:54
Install required MSYS2 image dependencies and direct gRPC and llama.cpp CMake builds to the bundled static zlib to avoid conflicts with the system DLL import library.

Signed-off-by: Lionel Colaso <lionelcolaso@outlook.com>
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch from 14e361d to 5edc684 Compare October 11, 2026 10:16

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Native windows version?

6 participants